Thursday, January 18, 2007

Science Blogging Conference

The North Caroline Science Blogging Conference starts this weekend. It looks like it covers some introductory how-to material, and then dives into issues such as promoting public understanding of science, teaching, and how does blogging interact with research.


If you're put off by the term 'blog' being overused and overhyped, think of the conference as discussing the potential and challenges of a (relatively) new one-to-many and many-to-many communication mechanism and how that could affect the interactions between scientists and the public, and the interactions between scientists.

Wednesday, January 17, 2007

Basis sets galore

More basis sets that you can shake a stick at over at the EMSL Basis Set Exchange. (Not sure why you would want to shake a stick at basis sets, though.)

Thursday, October 26, 2006

Schrodinger equation in the mass media

The Schrodinger equation appears in the background in the second half of Weird Al's video White & Nerdy.


The video also shows rolling dice (briefly), if one wishes to stretch its relevance to include Monte Carlo as well.

Friday, October 20, 2006

The Other QMC

Quasi-Monte Carlo is the other method commonly referred to by the QMC acronym (I will abbreviate it QuasiMC to minimize confusion). It's not really Monte Carlo since the sequence of points is not random, but the point sets do have a Monte Carlo-ish feel about them.

And in fact, it does share the property of arbitrary termination with Monte Carlo integration (consider a grid - you have have use all the points at a particular spacing. Stopping half way across the region of integration is not going to yield good results).

The reason people are interested in QuasiMC is that its convergence is better than the 1/sqrt(N) of Monte Carlo. It does this by using point sets that are more evenly distributed than random points - there are fewer clumps of points or large gaps that appear in random sequences (Search for 'Low-discrepancy sequences'. Here's the Wikipedia entries for a technical discussion and some sequences).

QuasiMC point sets are usually generated from sequences where the low-discrepancy condition can be verified theoretically, but it is possible to use the intuitive approach of making the points maximally spread out (see this presentation for an example).

The following algorithm will generate a set of N points that gives better convergence than random (at least it worked on a simple integrand in one dimension.)

  1. compute some random points (>> N points)

  2. pick a starting point and put it in the set S.

  3. find the point that is the furthest from all existing points in S, and place that point in S.

  4. repeat the step 3 until N points are selected.



(I'm not saying this an efficient way to generate the points, just that it can be done.)

For addition information, see also chapter 7.7 in Numerical Recipes

Friday, September 01, 2006

QMC wiki

Check out the QMC wiki at http://www.qmcwiki.org

It looks like a promising resource for the QMC research community.

Tuesday, June 27, 2006

Optimal histogram bin width

Kevin Knuth wrote a paper about finding the optimal number of bins to represent data in a histogram (Optimal Data-Based Binning for Histograms). He starts from a piecewise constant density model and finds the (Bayesian) posterior probability from this model (equation 36, which is actually the log of the posterior). The posterior function is then maximized to find the number of bins that best models the data.


The article also investigates the number of data points for a reliable estimation of the density. The recommendation is 100-150 points, if the distribution is Gaussian.


It would be interesting to apply this method to radial distribution functions. However the assumption of a constant volume for each bin is not met in this case. There are several ways this could be adjusted, but I'm not sure they are valid (scale each bin count by the volume, or use non-uniform bin spacing to maintain constant volume)


Alternately, the discussion references other algorithms for dealing with variable bin-width models (which may be better for resolving multiple peaks anyway).

Thursday, June 08, 2006

QMC derivation notes

I posted a document I wrote in grad school, Notes on the wavefunction and local energy. It contains derivations of various QMC formulas, particularly the first and second derivatives for several forms of wavefunctions.


I'm posting this for two reasons. The first is in case anyone finds the formulas useful when working on a QMC code.


The second is related to the process of scientific programming. When writing a QMC code, I found it useful to record the formulas and derivations in a neatly typset form. Then the next step involved turning the equations into computer code. (Then, of course, testing and debugging).


This workflow is what I would like to capture with the Progamming in Mathematical Notation work. The document with derivations could be written in content MathML (or something more amenable to human manipulation). Ideally the computer could then assist with verifying the derivations for correctness, and with converting the equations into computer code.


And as long as I'm dreaming, I'd really like a wiki-like interface for creating and editing such a document (making a set of hyperlinked pages rather than a single linear document)